Back

Briefings in Bioinformatics

Oxford University Press (OUP)

Preprints posted in the last 30 days, ranked by how well they match Briefings in Bioinformatics's content profile, based on 354 papers previously published here. The average preprint has a 0.32% match score for this journal, so anything above that is already an above-average fit.

1
HDOCK-Multimer: integrating docking and combinatorial assembly for structure prediction of large protein complexes

Yao, X.; Ya, Y.; Li, H.; Huang, S.-Y.

2026-08-06 bioinformatics 10.64898/2026.08.06.736029 medRxiv
Top 0.1%
21.9%
Show abstract

Deep learning methods, such as AlphaFold and RosettaFold, achieve high accuracy in protein structure prediction. However, predicting the structure of large protein complexes remains challenging due to their large size and intricate multi-chain interactions. Docking-based methods can handle large proteins, but are limited by the huge combinatorial binding space of multichains. Assembly-based approaches offer an alternative, but their accuracy critically relies on the precision of predicted subcomponents. Addressing the challenges, we propose HDOCK-Multimer (HDM), a structure prediction framework of large protein complexes by integrating ab initio docking and combinatorial assembly. HDM can efficiently reduce reliance on subcom-ponent accuracy through docking process, while leveraging the pairwise interactions of subcom-ponents through assembly strategy. HDM is extensively validated on three benchmarks of 35 large heteromeric complexes, 172 large protein complexes, and 7 CASP15 targets, and compared with state-of-the-art methods including MoLPC, CombFold, AlphaFold-Multimer (AFM), and AlphaFold3 (AF3). It is shown that HDOCK-Multimer substantially outperforms the other methods. In addition, HDM also shows ability to predict the stoichiometry and model the complex without stoichiometry input. It is anticipated that HDM will serve as a powerful tool for studying large protein complexes or molecular machines. The HDM package is freely available at https://github.com/huang-laboratory/HDOCK-Multimer/.

2
Predicting Protein-RNA Binding Affinity Changes via Spatial Coupling-Aware State Space Modeling

Chen, R.; Huang, X.; Jiang, H.; Ma, W.; Bi, X.; Wei, Z.; Nie, J.; Zhang, S.

2026-08-24 bioinformatics 10.64898/2026.08.23.745486 medRxiv
Top 0.1%
18.7%
Show abstract

Accurately predicting the effects of mutations on protein-RNA binding is crucial for elucidating disease mechanisms. Yet, exhaustively exploring the space of all possible variants is prohibitively expensive, motivating computational methods that can quantify mutation-induced changes in binding affinity (aka {Delta}{Delta}G) accurately and efficiently. We present iSCALE, an interpretable and generalizable deep learning method that adopts an implicit Spatial Coupling-Aware Ligand Encoding strategy to predict mutation-induced binding affinity changes. By injecting this implicit multiscale encoding scheme into a bidirectional state space modeling architecture, iSCALE learns a generalizable multiscale coupling pattern that achieves superior performances on not only the protein-RNA binding {Delta}{Delta}G, but also the protein stability {Delta}{Delta}G and protein-protein binding {Delta}{Delta}G predictions. Detailed analyses demonstrate that the model attention scores align well with structural characteristics. In addition, iSCALE shows good discriminative ability when predicting close samples such as complexes of same mutation but with different ligands or the same complex but with different mutation sites. In summary, iSCALE serves as an effective in silico tool for large-scale protein-RNA binding {Delta}{Delta}G prediction, which pushes the border of understanding in mutation-induced pathological outcomes.

3
scFair: Geometry-Aware Gene Budgets and Same-Rank Extension for Highly Variable Gene Selection

Li, Z.; James, A.; Li, S.

2026-08-14 bioinformatics 10.64898/2026.08.08.743679 medRxiv
Top 0.1%
18.6%
Show abstract

BackgroundHighly variable gene (HVG) selection begins almost every single-cell RNA-seq analysis. While ranking formulas have been compared extensively, the integer gene budget at which any ranking must be truncated is typically left to the user and habitually fixed near 2,000. Relying on such a convention carries hidden costs--lists that are too short erase subtle structure, whereas lists that are too long add noise and computational overhead. Moreover, because global rankings measure variance across all cells, markers for rare populations often lose the "variance vote count" to dominant bulk variation, leading to an unfair feature allocation at the hard cutoff. Whether this convention is defensible, and whether the budget and tail can be set from data without disturbing the ranking, has not been examined systematically. ResultsUnder a frozen seurat_v3 ranking, k-sweeps across 18 labeled datasets show that n = 2,000 is ARI-optimal on 1 of 18 datasets and that the best available budget is worth a mean ARI gain of +0.033 over it, establishing cardinality as a real and largely unexploited design axis. We present scFair, a Scanpy-compatible HVG layer that automates list length alone: geometry-aware auto_n sets a base size k from multi-seed density and stability features of an intermediate embedding (trading a modest, intentional compute increase for a safer data-driven default), and a same-rank append step acts as a conservative safeguard against cutoff unfairness by adding a short near-miss tail. The ranking is never recomputed or reweighted. On the 18-dataset panel, the default path improved Leiden-label agreement over HVG@2000 (median {Delta}ARI = +0.016; 13/5; Wilcoxon P = 0.0077) and outperformed the neighborhood-based selector triku at author defaults on 15/18 datasets (median +0.024; P = 0.004), while triku did not improve on HVG@2000. Controls locate the effect: cell-number-only rules do not beat HVG@2000, an FDR-chosen length imposed on the frozen ranking is flat, and a fixed HVG@2200 default is not a general substitute because it cannot produce the short lists that compact matrices call for. ConclusionsA fixed budget near 2,000 HVGs is frequently suboptimal, and list cardinality is a separable design axis that can be automated without changing the ranking formula. Effect sizes are modest, the short-list branch rests on four datasets, and rule thresholds were developed with partial overlap to the evaluation panel.

4
Causally-inspired meta-representation learning framework for predicting patient-specific clinical responses to drug combinations

Zhang, Q.-Q.; Zhang, S.-W.; Shi, M.-H.; Li, J.-N.; Qiang, Y.-R.; Zhang, T.-H.

2026-08-21 bioinformatics 10.64898/2026.08.13.744613 medRxiv
Top 0.2%
18.6%
Show abstract

Large-scale prediction and assessment of clinical patient responses (i.e., RECIST class) to drug combinations remains challenging due to scarce patient-derived data. The existing prediction methods mainly rely on cancer cell line models. However, substantial biological heterogeneity between cancer cell lines and cancer patients within same tissues, as well as the heterogeneity between one tissue and another, often limit the generalizability of these methods in clinical patients. To overcome these limitations, here we present CaMeRe, a Causally-inspired Meta-representation learning framework designed to predict patient-specific clinical Response to drug combinations. In situations where stable causal factors and domain-specific response-modulating factors are unobservable, explicit discrete domain labels are unavailable, and data is scarce, CaMeRe designed a domain-invariant causal representation learning (DICRL) model guided by the invariant information bottleneck theory and causal intervention invariance principle, and also built a meta-learning framework with bi-level domain generalization to optimize DICRL model for achieving multi-domain generalization within and across tissues. By integrating the causal representation learning and meta learning framework, CaMeRe not only exhibited robust multi-domain generalization performance across multiple clinical drug combination response datasets and PDXs drug combination response datasets and generalization scenarios, but also had better interpretability. We applied CaMeRe to predict drug-combination response scores for 3,423 patients across 542,080 drug combinations. The predicted scores were significantly associated with biomarkers of known drug combinations and enabled the prioritization of candidate drug combinations across 11 cancer types, with stronger support from literature and clinical trial evidences than random baselines. We believe that CaMeRe can be a useful tool for predicting large-scale clinical individual drug combination responses and it has broad clinical applications.

5
Benchmarking single-cell foundation models in a zero-shot setting

Gaballa, Y.; Ahmed, S.; Abdelaal, T.

2026-08-07 bioinformatics 10.64898/2026.08.03.739553 medRxiv
Top 0.2%
18.3%
Show abstract

Single-cell foundation models have recently emerged as a promising approach for learning general- purpose representations from large-scale transcriptomic data. These models are trained on millions of cells and are designed to transfer their learned representations to a wide range of downstream tasks. However, their practical benefits compared to traditional approaches are still not fully understood. This study evaluates four foundation models, namely scGPT, SCimilarity, UCE, and Transcriptformer, across four downstream tasks: cell type annotation, human data integration, cross-species data integration, and protein expression prediction. Embeddings generated by each model were assessed using multiple public single-cell datasets and compared against conventional machine learning baselines. Performance was measured using task-specific evaluation metrics, including classification, integration, and regression metrics. The results showed that foundation model embeddings did not consistently outperform traditional approaches. In the cell type annotation task, baseline methods achieved the strongest performance across most datasets. For protein expression prediction, however, embeddings from the foundation models generally produced more accurate predictions than the baseline, with SCimilarity achieving the lowest prediction error and Transcriptformer obtaining the highest correlation scores. In the data integration task, all foundation models produced moderate results, while scVI (the baseline) achieved the strongest integration performance. Overall, the results suggest that current single-cell foundation models provide useful representations for some downstream tasks in zero-shot conditions but do not yet offer a universal replacement for task-specific methods. Their effectiveness remains dependent on the application and evaluation setting.

6
A distribution-aware and functionally relevant novel framework for generation and discovery of bioactive peptides

Abhigyan, R.; Sood, V.; Arora, P.; Kaur, B.

2026-08-09 bioinformatics 10.64898/2026.08.04.742799 medRxiv
Top 0.2%
15.6%
Show abstract

Recent advances in artificial intelligence have accelerated the discovery of bioactive peptides by enabling computational exploration of the vast peptide sequence space. However, existing peptide generation approaches generally rely on either distribution-learning models, which generate biologically realistic sequences but do not consistently optimize functional activity, or optimization-based methods, which maximize prediction confidence while often deviating from the underlying distribution of experimentally validated peptides. To address this limitation, a two-phase generative-evolutionary framework is proposed that integrates distribution learning with evolutionary optimization. In the first phase, Variational Autoencoders (VAE), Autoregressive Transformers (ART), and Token Diffusion Transformers (TDT) are used to generate biologically plausible seed peptides. In the second phase, these peptides were used as initial seed for Hill Climbing optimization procedure that iteratively improves fitness function score. The proposed two-phase framework was evaluated using a dataset of experimentally validated IL-2-inducing peptides. Evaluation using independent IL-2 prediction models showed that Autoregressive Transformer combined with Hill Climbing achieved the best overall performance, achieving the mean IL-2 induction confidence score of 0.96 while reducing KL divergence from 2.26 for standalone Hill Climbing to 0.75. A case study on an independent IL-13 inducing peptide dataset showed similar trends, with ART initialized Hill Climbing achieving the mean IL-13 induction score of 0.99 while reducing KL divergence from 1.76 to 0.59. Overall, the framework provides a generalizable approach for balancing functional optimization and distributional realism and can be applied to peptide discovery and data augmentation in imbalanced biological datasets thereby generating high confidence peptides for wet lab validation. HighlightsO_LIProposed a two-phase framework for bioactive peptide generation with potential to address class imbalance in peptide classification tasks. C_LIO_LIPerformed a systematic comparison of distribution-learning and optimization-based approaches for peptide generation. C_LIO_LICombined distribution-learning models for sequence generation with optimization algorithms for improving peptide functional properties. C_LIO_LIDemonstrated the applicability of the proposed framework across multiple bioactive peptide datasets. C_LI

7
SAMP V2: A novel stacking ensemble learning model for antimicrobial peptides identification based on augmented split amino acid composition with biochemical-sequence-order information

Sun, M.; Wang, J.; Wan, S.

2026-08-19 bioinformatics 10.64898/2026.08.12.744552 medRxiv
Top 0.2%
15.2%
Show abstract

Antimicrobial resistance reduces the effectiveness of conventional antibiotics and has become a major global health threat, highlighting the need for new anti-infective agents. Antimicrobial peptides (AMPs), a diverse class of innate immune effectors with broad-spectrum antimicrobial activity, are promising candidates for combating drug-resistant infections. Identifying AMPs by wet-lab experiments, however, remains costly and time-consuming, creating a strong demand for computational identification methods. Our recently developed method, SAMP, captures region-specific residue distributions based on proportionalized split amino acid composition. However, SAMP might ignore key biochemical information and sequence order information. Here we present SAMP V2, a stacking ensemble learning framework based on biochemical and sequence-order information augmented split amino acid composition (BIA-SAAC), which extends SAMP by integrating pseudo-amino acid composition features with biochemical and sequence-order information into split peptide regions. Specifically, each peptide is divided into N-terminal, middle, and C-terminal regions, and pseudo amino acid composition is calculated within each region. Benchmarking tests on six independent test datasets, SAMP V2 outperformed multiple state-of-the-art models, including AMPpred-MFA and iAMP-Attenpred, in terms of accuracy, MCC, G-measure and F1-score. Given its high and robust performance, SAMP V2 could significantly accelerate the discovery of next-generation antimicrobial therapeutics for addressing the global threat of multidrug-resistant pathogens.

8
Cross-attention and language models reveal the interpretability of functional predictions for the human olfactory receptor family

Zhang, Y.-F.; Xu, Z.-h.; Gao, C.-x.; Duan, S.-Y.; Li, G.; Xu, C.; Lu, H.-M.

2026-08-18 bioinformatics 10.64898/2026.08.10.744067 medRxiv
Top 0.3%
13.0%
Show abstract

The attention mechanism offers the possibility for data-driven discovery of biological principles. However, for important protein families such as human olfactory receptors, the extent to which attention can associate with biologically meaningful key regions lacks systematic validation. In this study, using human olfactory receptors (ORs) as a model, we constructed CrossVOI, a VOC-OR interaction prediction framework based on protein language models and cross-attention, achieving predictive performance superior to existing methods. Furthermore, we systematically analyzed the attention distributions of CrossVOI and found that attention not only focused on ligand-binding interfaces and evolutionarily conserved sites, but also to some extent identified certain dynamically regulated regions. In summary, we propose CrossVOI, currently the best-performing framework for VOC-OR interaction prediction, and analyze the interpretability of the attention mechanism for human ORs. This study provides insights into the interpretability of protein function prediction methods and is expected to contribute to the exploration of attention mechanisms in biological mechanisms, and provide assistance for large-scale screening and mechanistic analysis of olfactory receptors.

9
Correction of the cytosine deamination artifacts in FFPE-based sequencing experiments

Płonka, W.; Kostka, D.; Lalik, A.; Kurpas, M.; Dinh, K. N.; Sitkiewicz, M.; Kimmel, M.; Rzyman, W.; Jaksik, R.

2026-08-19 bioinformatics 10.64898/2026.08.11.744151 medRxiv
Top 0.3%
13.0%
Show abstract

Formalin-fixed, paraffin-embedded (FFPE) tissues remain an essential resource for molecular studies, yet formalin-induced cytosine deamination introduces characteristic C>T/G>A artifacts that compromise the accuracy of next-generation sequencing (NGS) analyses. Numerous computational methods and enzymatic DNA repair strategies have been proposed to reduce these artifacts, but no systematic comparison across tools and experimental conditions exists. Here, we evaluate the performance of seven computational approaches (SOBDetector, Ideafix, MicroSEC, FFPolish, DeepOmics FFPE/FFPE-PLUS, FFPErase) together with the NEBNext(R) FFPE DNA Repair Mix v2, a multi-enzyme repair system applied during DNA preparation. Using three independent datasets, one based on whole genome sequencing (CGCI-BL) and two on whole exome sequencing (TCGA-PC and SUT-LUAD, the latter containing enzymatically repaired samples), and matched fresh-frozen samples as the gold standard, we assess precision, sensitivity, and artifact reduction efficiency across all methods. We further examine the potential synergy between enzymatic repair and post-sequencing computational filtering. Our results provide practical guidelines for FFPE artifact correction and demonstrate that enzymatic treatment provides the best results, while among the computational methods, FFPErase offers the most robust reduction of cytosine deamination artifacts while maximizing the retention of true somatic variants. KEY MESSAGESO_LIFormalin fixation in FFPE samples introduces artifacts that can significantly affect the accuracy of NGS analyses. C_LIO_LIAmong the evaluated approaches, enzymatic repair using NEBNext(R) FFPE DNA Repair Mix v2 achieves the most effective reduction of sequencing artifacts. C_LIO_LIComputational methods vary in performance, with FFPErase showing the most robust balance between artifact removal and retention of true somatic variants. C_LIO_LICombining enzymatic repair with computational filtering did not lead to consistent improvements in performance across datasets. C_LI

10
Benchmarking Imputation Methods for Single-Cell RNA Sequencing Data Using Peripheral Blood Mononuclear Cells from Acute Myocardial Infarction Patients

Ramesh, P.; Fyta, M.

2026-08-27 bioinformatics 10.64898/2026.08.23.746230 medRxiv
Top 0.4%
12.1%
Show abstract

Acute myocardial infarction (AMI) remains one of the leading causes of mortality worldwide, and the following post-effects, such as post-AMI inflammation and tissue repair, involve peripheral blood mononuclear cells playing a critical role. The influence of imputation methods in biological data is assessed with respect to high-resolution single-cell RNA sequencing (scRNAseq) data relevant to these cells. Still scRNAseq data often encounter a lot of dropout events, leading to sparse and noisy datasets, hampering downstream results. To assess the influence of the missingness in the data, we artificially impose different levels of dropout in available scRNAseq data by leveraging various imputation techniques. Specifically, we introduce artificial missingness at 10%, 20%, and 30% levels under a missing completely at random (MCAR) framework, repeated across 10 independent runs. We benchmarked six imputation strategies - MAGIC, IterativeImputer, KNNImputer, Mean Imputation, SoftImpute, and a Generative adversarial network (GAN) - based approaches using multiple evaluation metrics: marker gene preservation, clustering consistency (Adjusted Rand Index - ARI), gene-wise correlation with ground truth, and structural separation (silhouette scores). The results clearly underline that no single imputation method dominated across all metrics. Overall, Mean and KNN imputers showed limited recovery across all benchmarks. GAN excelled in global transcriptional recovery and SoftImpute in preserving biologically meaningful cell-type signals. Our results highlight the importance of selecting the imputation methods as part of the pre-processing step towards the downstream biological questions related to transcriptome recovery, detection of marker genes, or maintaining cell-type-specific resolution.

11
What Makes a Good Vaccine Antigen Target? Defining Key Features and Predicting Candidates in the Staphylococcus aureus Proteome

Prasetyo, N. K.; Langley, R. J.; Radcliff, F. J.; Gardner, P. P.

2026-08-10 bioinformatics 10.64898/2026.08.09.743793 medRxiv
Top 0.4%
12.0%
Show abstract

The rapid advancement of computational methods is transforming vaccine development by enabling faster, data-driven identification of promising antigens. In this study, we applied an in-silico pipeline to assess a broad set of sequence, structure, localisation, and immunology-derived features and determine which most effectively discriminate antigens from non-antigens in bacteria. Using these insights, we identified bacterial proteins with high potential as vaccine antigens. Applied to Staphylococcus aureus, this approach prioritized 304 candidate antigens, highlighting SSLs, nutrient acquisition factors, and cell wall-associated enzymes. While these findings demonstrate the potential of bioinformatics-guided antigen discovery, experimental validation remains essential. This work underscores the growing role of integrated computational and machine-learning approaches in accelerating next-generation vaccine design.

12
NACraft: Programmatic nucleic-acid aptamer design via all-atom structure-model feedback

Zhu, H.; Wang, J.; Zhao, W.; Xu, Y.; Su, H.; Wang, J.; Wang, Q.; Yu, Y.; You, Z.; Du, G.; Heng, P. A.; Zhang, L.; Zhang, O.

2026-08-21 bioinformatics 10.64898/2026.08.15.744087 medRxiv
Top 0.4%
11.9%
Show abstract

Protein-nucleic-acid interactions underpin diverse biological processes and provide a basis for molecular sensing, regulation and therapeutic intervention. However, the coupled dependence of aptamer function on nucleotide sequence, three-dimensional folding and target binding makes rational RNA and DNA binder design challenging. Here we present NACraft, a training-free and programmatic framework for all-atom nucleic-acid aptamer design based on backpropagation through structure-model feedback. By composing binding, sequence-similarity and anti-binding constraints, NACraft supports de novo generation, similarity-guided sampling and target-selective design within a unified optimization framework, without task-specific training or fine-tuning. Computational experiments showed that NACraft generated high-confidence candidates de novo across diverse protein targets, with further improvements achieved through similarity-guided design for both RNA and DNA complexes. Its target-selective design capability was further validated in silico, with 69.44% of paired candidates generated to favour the positive target EGFR over the off-target HER2. Under matched independent AlphaFold3 evaluation, NACraft achieved better performance than ODesign in 10 of 11 NA-12 targets and 17 of 20 protein target-length settings. Together, these results demonstrate the effectiveness and versatility of NACraft and extend structure-model hallucination toward programmatic nucleic-acid aptamer design. Codehttps://github.com/OTEAM-AI4S/NACraft

13
A Two-Stage ESM-Based Machine Learning Pipeline for Robust Hierarchical Enzyme Function Prediction

Hua, X.; Grimaud, G. M.

2026-08-18 bioinformatics 10.64898/2026.08.14.744831 medRxiv
Top 0.5%
11.5%
Show abstract

Accurate enzyme annotation remains a major bottleneck in translating rapidly growing protein sequence data into biological knowledge. Enzyme Commission (EC) prediction is particularly challenging because enzyme functions are organized hierarchically, annotations are often imbalanced across classes, and sequence similarity alone may be insufficient to resolve functional differences. To address these challenges, we developed ESM-ECForest, a two-stage framework that combines protein embeddings generated by the pretrained language model ESM-2 (Evolutionary Scale Modeling 2) with Random Forest classifiers. The first stage distinguishes enzymes from non-enzymes, whereas the second assigns one or more EC numbers to proteins predicted to be enzymatic. On an external benchmark comprising 25,778 protein sequences, ESM-ECForest achieved the highest weighted F1 score among the evaluated methods at all four EC levels, decreasing from 0.94 at Level 1 to 0.90 at Level 4. The largest relative improvements were observed for lyases (EC 4), ligases (EC 6), and translocases (EC 7), although EC 6 and EC 7 remained the most difficult classes internally. Visualization of the ESM-2 embedding space using Uniform Manifold Approximation and Projection (UMAP) revealed clustering patterns consistent with enzyme functional relationships, indicating that biologically relevant information is retained in the pretrained representations prior to supervised classification. These results support the use of pretrained protein language model embeddings as an effective foundation for enzyme annotation. By combining large-scale sequence representations with a lightweight supervised classifier, ESM-ECForest provides a scalable approach for EC prediction and may facilitate functional annotation of protein sequences derived from large genomic and metagenomic datasets.

14
Data-Centric Evaluation of Protein Function Prediction Pipelines

Soto-Garcia, N.; Murillo-Acevedo, N.; Garcia Vinuesa, J.; Islas-Avila, A. L.; D. Davari, M.; Murgas, L.; Hassanin, A.; Orostica, K.; Gonzalez-Puelma, J.; Navarrete, M.; Rebollar-Martinez, A.; Uribe-Paredes, R.; Cadet, F.; Medina-Ortiz, D.

2026-08-11 bioinformatics 10.64898/2026.08.05.742971 medRxiv
Top 0.5%
10.6%
Show abstract

Performance estimates in protein function prediction depend not only on model choice but also on upstream decisions that define the learning problem. Using antioxidant protein classification as a controlled case study, we evaluated how dataset harmonisation, protein representation, redundancy control, and partitioning strategy affect protein machine learning pipelines. We integrated 18,804 records from 12 publicly available dataset entries into a curated consensus dataset of 4,193 protein sequences. One-hot encoding and six pretrained protein language model representations were evaluated as model inputs and as similarity spaces for redundancy reduction and distance-aware splitting. Representation choice substantially altered dataset geometry, retained dataset size, class balance, and downstream evaluation. At representation-specific p90 thresholds, one-hot encoding retained the complete dataset, whereas pretrained embeddings retained between 5% and 25% of sequences. Distance-aware partitioning reduced apparent performance relative to random splitting by up to 0.15 MCC before redundancy control, while this difference narrowed after similarity filtering. Selected configurations nevertheless maintained high performance under stricter evaluation, reaching an MCC of 0.84. These findings show that performance estimates should be interpreted as outcomes of complete data-centric workflows rather than isolated properties of predictive models.

15
Detecting CYP2C19 deletions from genotyping array signals using neural networks

Yelmen, B.; Hofmeister, R. J.; Lutsar, V. K.; Finianos, M.; Stone, B. C.; Joeloo, M.; Krebs, K.; Kivistik, P. A.; Smit, S.; Estonian Biobank Research Team, ; Metspalu, M.; Hudjashov, G.; Milani, L.

2026-08-25 bioinformatics 10.64898/2026.08.21.746170 medRxiv
Top 0.6%
9.9%
Show abstract

Since copy number variations (CNVs) in pharmacogenes can cause significant alterations in drug metabolism, their reliable detection is of high importance both for large-scale studies and personalized medicine. Whole-genome sequencing, and specifically long-read sequencing, is the gold standard for CNV detection. Despite increasing availability of these technologies, genotyping arrays are still widely used as cost-effective alternatives in biobank and clinical settings, yet calling CNVs based on array intensity signals is challenging due to low base pair resolution. In this work, we developed a neural network model, nnCNV, to predict deletions in the CYP2C19 pharmacogene region from array intensity signals. We compared our method to the most widely used algorithm, PennCNV, and demonstrated better performance reaching 100% accuracy in the test dataset. Furthermore, we predicted probe-by-probe CYP2C19 deletion coordinates for all Estonian Biobank samples using nnCNV and PennCNV, and validated these predictions using an identity-by-descent (IBD) sharing method, which also demonstrated superior nnCNV performance. For the deletion samples with conflicting PennCNV and nnCNV predictions, we performed PCR analysis for validation, which showed 97% precision for nnCNV compared to 23% for PennCNV. Finally, we assessed the gradient-based feature importance maps and showed that nnCNV utilizes signal intensity information not only from deletion probes, but also from probes in flanking regions. Our results demonstrate that long-range information, which cannot be utilized by hidden Markov models, can improve CNV calling.

16
Ori-Finder-Arch: An Updated Web Server for the Annotation and Visualization of Archaeal Replication Origins

You, Z.; Zhang, Z.; Luo, H.; Gao, F.

2026-08-19 bioinformatics 10.64898/2026.08.15.744077 medRxiv
Top 0.6%
9.9%
Show abstract

Archaea are promising chassis organisms in biotechnology, and the accurate annotation of their chromosomal replication origins (oriCs) is the key to unlocking their full potential. However, the existing Ori-Finder 2 web server suffers from low accuracy, slow speed, and limited scalability. In this study, we present Ori-Finder-Arch, an updated web server for high-performance oriC prediction in archaea. This pipeline integrates HMMER-based replication initiation protein (RIP) annotation, refined consensus motif recognition, and GC profile-based DNA unwinding element (DUE) detection. On a benchmark set of experimentally validated oriCs, Ori-Finder-Arch achieved a recall of 95.6% and a precision of 86.0%, substantially outperforming Ori-Finder 2 (62.2% and 63.6%, respectively), while running 4.75 times faster and supporting diverse assembly levels. When applied to the available archaeal assemblies, it successfully annotated 17,472 oriCs. Meanwhile, the web server provides interactive visualizations at different levels. In conclusion, Ori-Finder-Arch offers an efficient, accurate, and user-friendly platform for advanced studies of archaeal DNA replication initiation and synthetic biology applications, and is freely available at https://tubic.org/Ori-Finder-Arch/ and https://tubic.tju.edu.cn/Ori-Finder-Arch/.

17
Learning Shared Residue Backgrounds and Modification-Specific Offsets for PTM Site Prediction

Pokharel, S.; Bhusal, B.

2026-08-11 bioinformatics 10.64898/2026.08.05.743079 medRxiv
Top 0.6%
9.7%
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWPost-translational modifications (PTMs) are chemical changes added to proteins after translation. These changes affect protein function and regulation, and their disruption is linked to disease-associated mechanisms. Because experimentally validating all possible modification sites is impractical, many computational predictors have been developed for PTM site prediction. In this work, we study whether a shared model can represent common residue-background patterns while learning modification-specific background-to-positive offsets. This framing is especially relevant for residues such as lysine (K), which can be acetylated, ubiquitinated, methylated, or sumoylated depending on the surrounding protein context. We propose an anchor-guided rectified flow matching framework for multi-type PTM site prediction from protein language model embeddings. For each PTM-residue pair, the model builds residue-background anchors from PTM-compatible unannotated residues and positive anchors from experimentally annotated modified residues. Given a candidate residue and target modification type, the model compares the residue embedding with these anchor sets and uses a rectified flow module to estimate a modification-conditioned background-to-positive offset. This offset is combined with anchor-based features and used for site scoring. We evaluate the framework on a dbPTM-derived benchmark covering six commonly studied PTMs: phosphorylation, acetylation, ubiquitination, methylation, sumoylation, and N-linked glycosylation. In the shared-model setting, our approach achieves a macro AUPRC of 0.4195, improving over the gated multi-anchor baseline of 0.4154, while independently trained per-modification models achieve 0.4353. These results suggest that multi-type PTM prediction can be modeled within a single shared framework by combining residue-background anchors with modification-conditioned offset features.

18
A Principled Statistical Framework for Analyzing Spatial Patterns in Spatially Resolved Multi-Omics

Li, J.; Raina, M.; Wang, Y.; Zeng, S.; Yu, Y.; Yu, X.; Jin, X.; Chang, Y.; Feliciano, D.; Himmelfarb, J.; Ricardo, A. C.; Nachman, P. H.; Vazquez, M.; Caramori, M. L.; Barisoni, L.; Kretzler, M.; Jain, S.; Dagher, P. C.; El-Achkar, T. M.; Eadon, M. T.; Human Biomolecular Atlas Program, ; Kidney Precision Medicine Project, ; Melo Ferreira, R.; Ma, Q.; Wang, J.; Xu, D.

2026-08-10 bioinformatics 10.64898/2026.08.04.742894 medRxiv
Top 0.6%
9.7%
Show abstract

Emerging spatial multi-omics technologies enable the profiling of molecular variation within its tissue context, yet existing methods for identifying spatially variable features lack principled approaches to experimental design and cross-sample inference. Here, we present STORM, a principled Statistical TOol for spatially Resolved Multi-omics, for rigorously analyzing spatial patterns in spatial multi-omics research. STORM incorporates a robust and efficient nonparametric test that quantifies local deviations in molecular feature measurements to detect spatial dependence across transcriptomic and proteomic data. It further estimates an interpretable spatial effect size, supports power calculations for both spatial locations and biological replicates, and enables formal group-level comparisons. In several simulated and experimental spatial multi-omics case studies, STORM demonstrates reliable performance in detecting spatial structures while offering quantitative support for study design decisions. Overall, STORM provides a principled statistical framework that unifies spatial hypothesis testing, effect size estimation, power analysis, and experimental design for spatial multi-omics data.

19
PMPNN-DDG: an accurate machine learning-based {triangleup}{triangleup}G prediction pipeline trained on a novel interpretable feature set extracted from ProteinMPNN

Jani, R.; Ahmed, S.

2026-08-27 bioinformatics 10.64898/2026.08.23.746499 medRxiv
Top 0.6%
9.7%
Show abstract

An accurate and tractable approximation of the single-point mutation-induced change in protein thermodynamic stability, denoted by DDG, is critical for understanding the genotype-phenotype relationship. Several computational methods have been proposed for this problem; however, limited and error-prone training data and the difficult-to-predict magnitude of structural perturbations make this a challenging task. Consequently, the computational predictors proposed throughout the past decade incrementally improved prediction performance by proposing novel features, combining existing features, task-adapted neural network architectures, loss functions, data augmentation techniques, and pre-training procedures. In this work, we propose PMPNN-DDG, a Random Forest-based DDG prediction model, trained on a novel set of interpretable features extracted from the recently proposed message-passing neural network-based fixed backbone protein design model, ProteinMPNN. On the S669 independent test set, PMPNN-DDG achieves rF +R = 0.64 and RMSE = 1.45, outperforming all compared baseline methods across the reported evaluation measures. On the Ssym independent test set, it achieves rF +R = 0.81, rF -R = -0.99, and RMSE = 1.10, showing competitive performance relative to the compared baselines. PMPNN-DDG is publicly available at https://github.com/dRanger666/PMPNN-DDG.

20
Evaluating Lightweight and Full Fine-Tuning Strategies Against Classical Machine Learning for Protein Function Prediction

Ab Ghani, N. S.; Matsushita, T.; Noguchi, T.; Kurumida, Y.; Kawada, S.; Ito, T.; Umetsu, M.; Saito, Y.

2026-08-07 bioinformatics 10.64898/2026.08.02.737389 medRxiv
Top 0.7%
9.6%
Show abstract

Motivation Protein language models (PLMs) have emerged as powerful tools for sequence-based prediction of protein function, yet systematic benchmarks comparing frozen embeddings, fine-tuning strategies like Low-Rank Adaptation (LoRA) and classical machine learning (ML) remain limited. We benchmarked four ML strategies: ML using amino acid descriptors (SL-AAFeat), ML using frozen embeddings from 20 PLMs across various pooling strategies (SL-Embed), full model fine-tuning (FT-Full) and LoRA-based fine-tuning (FT-LoRA). Performance was evaluated on the in-house VHH phage display dataset (VHH) for binding affinity prediction and the TAPE fluorescence dataset (FLS and FLS10) for mutational effect prediction. Results Model performance depended strongly on the dataset and adaptation strategy. Max pooling consistently improved embedding-based models, while amino acid descriptors remained competitive under specific datasets and resource constraints. Fine-tuning generally provided the highest predictive performance, but the advantage is not universal. Hyperparameter optimization significantly enhanced FT-LoRA, enabling it to outperform FT-Full on the VHH dataset with less than 10% model parameter adaptation. In contrast, FT-Full achieved the best performance on FLS and FLS10. Several medium-sized PLMs performed comparably to larger models, highlighting favorable performance-efficiency trade-offs. Overall, this paper presents a thorough review of PLM utilization strategies and practical recommendations for selecting suitable strategies based on dataset characteristics and available computational resources. Availability The source code used in this manuscript is available in a Zenodo repository at https://doi.org/10.5281/zenodo.21466255.